National Science Review
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match National Science Review's content profile, based on 21 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Zhao, D.; Yang, Y.; Sun, J.; Zhang, J.; Duan, H.; Tan, Y.; Liu, l.
Show abstract
Although the "RNA world" hypothesis suggests that RNA played a crucial role in the origin of life [7], the functional framework of RNA in prebiotic protein synthesis and the mechanisms of genetic code formation during the prebiotic period remain poorly understood. Here, using the prebiotic "primordial soup" as a model, we reconstructed the detailed steps that would yield a protein with a stable ordered amino-acid sequence in the "primordial soup" at the prebiotic period. In the "primordial soup", a large number of medium- to large-sized biomolecule-like substances--such as RNA-like and protein-like molecules of various sizes and shapes, as well as related polymers like amino-acid-RNA-like etc.--did generate and accumulate. Moreover, protein-like and RNA-like molecules formed even more intricate complexes. These complexes bound free mRNA-like molecules through complementary base pairing. Subsequently, with an extremely low probability, two adjacent amino-acid-RNA-like molecules became bound to this free mRNA-like molecule, and their amino acids underwent a condensation reaction by the complexes, producing peptides and eventually proteins or polypeptides. This free mRNA-like molecule exhibits a certain flexible structure, whereas the super-large complexes formed by protein-like and RNA-like molecules (which possess certain activities) and the amino-acid-RNA molecules exhibit relatively rigid structures. Long-term evolution and mutual selection led to the emergence of proteins with stable amino acid sequences and moderate catalytic activity. In this way, the nucleotide information embedded in such mRNA-like molecules indirectly express through protein synthesis--a process we term the "A Co-Adaptation Flexible-Rigid Docking Model", where flexible mRNA-like molecules dock onto rigid complexes to enable ordered peptide formation. Finally, we show how trinucleotide codons emerge naturally from the flexible-rigid docking constraints.
Mishra, P.; Bhattacharya, S.; Bhattacharya, J.; Jain, Y.; Sandhu, K. S.
Show abstract
The naked mole rat is an evolutionary outlier among mammals, exhibiting extreme longevity, cancer resistance, hypoxia tolerance, pain insensitivity, eusociality, poikilothermy and other distinctive physiological traits, most of which likely resulted from its adaptation to highly adverse subterranean habitat. Despite accumulating data, the genetic and molecular basis underlying these traits remain poorly understood. Through analyses of 18 distinct protein attributes and allied datasets across hystricomorphs, myomorphs, carnivores, and primates, we observed lineage-specific evolutionary divergence in intrinsic protein disorder in the naked mole rat. The disorder turnover exhibited functional dichotomy. The gain of disorder preferentially associated with proteostasis, immune regulation, neurodevelopment, skeletal growth and tumour suppressive properties, while loss of disorder modulated mostly the cardiac development. The proteins that gained disorder in NMR exhibited lower degradation rates, consistent with stabilization through phase-separation, while the proteins losing disorder show pronounced divergence in gene expression. The disorder turnover was primarily driven by indels affecting functional regions including Pfam domains, ANCHOR-predicted binding sites, short linear motifs, stress induced modifications of Tyr, Met, and Cys residues. Notably, the gained disordered regions were inferred to be redox-sensitive, aligning to exceptional stress tolerance in naked mole rats. Collectively, our results highlight an unusual and previously overlooked large-scale proteome remodelling that drives the molecular evolution of extraordinary traits of naked mole rat. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=169 SRC="FIGDIR/small/720964v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@644d1dorg.highwire.dtl.DTLVardef@102c9e4org.highwire.dtl.DTLVardef@14c770org.highwire.dtl.DTLVardef@31bc5e_HPS_FORMAT_FIGEXP M_FIG C_FIG
Zuo, N.; Cai, X.; Wang, W.; Ren, Z.; Jiang, Z.; Jiang, W.; Song, X.; Gu, Y.
Show abstract
Nicotine accumulates in the gut and drives non-alcoholic steatohepatitis (NASH) via the gut-liver axis, yet no effective clinical intervention is currently available. To address this challenge, the probiotic Escherichia coli Nissle 1917 (EcN) was engineered for in situ nicotine clearance in the gut. Mutational screening of nicotine oxidoreductase 2 (PpNicA2) identified a highly active variant, PpNicA2A107R. Its incorporation into EcN together with an electron transfer protein (CycN) and a newly identified transporter (T3/T7) yielded 80% nicotine-degrading activity. Chromosomal integration of this module generated a stable strain, EcN-N12, which in NASH mouse models depleted intestinal nicotine, rescued hepatic lipid metabolism, alleviated tissue damage, and intercepted the nicotine-mediated gut-liver axis pathological progression. This work thus offers an effective and clinically translatable approach for nicotine-associated diseases.
Zhang, T.; Xiong, Y.; Chen, K.; Wu, S.; Yan, X.; Zhou, J.; Wang, Y.; Yang, C.; Wang, P.; Zhou, Z.
Show abstract
Camptothecin derivatives are first-line anticancer drugs used worldwide for the treatment of diverse malignant tumors. However, the biosynthetic pathway of camptothecin has remained elusive for five decades. Here, we fully map its entire biosynthetic route. We discovered five key missing enzymes (OpCAR, OpSDR11, OpCS, OpGH1, and OpSTR) via the combination of MALDI mass spectrometry imaging, single-cell RNA sequencing and co-expression analysis. Meanwhile, we demonstrated a free flavin mononucleotide triggered the non-enzymatic 6-5-6 to 6-6-5 fused-ring skeleton rearrangement, filling the last gap in camptothecin biosynthesis. Finally, we validated this identified pathway and achieved the de novo biosynthesis of camptothecin in Saccharomyces cerevisiae. These discoveries uncover the long-standing mystery underlying camptothecin and pave the way for manufacturing camptothecin and its derivatives through synthetic biology approaches.
Lin, G.; Miao, R.; Sacheck, J.; Zhang, X.
Show abstract
Physical activity (PA) plays an important role in maintaining and improving health. Daily steps have been a key PA measure that is easily accessible with common wearable devices. However, methods are lacking to recommend a personalized optimal distribution of daily steps over a period of time for the best of certain health biomarkers. In this paper, we fill this void based on the data from the All of Us Research Program which includes months of step counts as well as repeated measurements of key health biomarkers. We develop a new offline reinforcement learning (RL) algorithm to learn personalized and optimal PA distributions associated with cardiometabolic risk, where the action is a function representing the daily step distribution over a period of time. Simulation studies demonstrate the advantage of the proposed approach over existing continuous-action RL methods. The learned optimal policy from the All of Us data generally suggests people take more daily steps and also follow a more consistent pattern of PA over time while offering tailored recommendations for subgroups in blood glucose level, body mass index, blood pressure, age, and sex.
Liao, H.; Qin, B.; Zhou, L.
Show abstract
Objectives; The role of nuclear receptor subfamily 4, group A, member 3 (NR4A3) in hepatic steatosis, inflammation, and insulin resistance (IR) within the context of metabolic dysfunction-associated steatotic liver disease (MASLD) remains largely underexplored. Consequently, this study aimed to examine NR4A3's impact on MASLD and the potential underlying mechanisms. Methods; We aimed to elucidate the functional role of NR4A3 in MASLD through its knockdown in cell culture and animal models. To establish the cell culture model of MASLD, LO2 cells were treated with free fatty acids (FFAs), while male C57BL/6 mice were fed a high-fat diet (HFD) to create the animal model. NR4A3 knockdown was achieved using specific short hairpin RNA (NR4A3-shRNA) in the mice model and three small interfering RNAs (NR4A3-siRNAs) in the cell culture model. The lipids content, fatty acid synthesis, inflammatory factors, and IR were then assessed with and without NR4A3 knockdown. Furthermore, the underlying mechanism through which NR4A3 exerts its influence was explored by analyzing the interaction between NR4A3 and activating transcription factor 3 (ATF3). Results: In the cell culture experiments, the knockdown of NR4A3 significantly decreased the lipids content, fatty acid synthesis, and inflammatory factors in the LO2 cells treated with FFAs in the NR4A3-shRNA group compared with those in the NC-shRNA control group. In the animal model experiments, NR4A3 knockdown in the HFD male C57BL/6 mice significantly ameliorated HFD-induced hepatic steatosis, inflammation, and IR. Mechanistically, the knockdown of NR4A3 downregulated the expression and transcriptional activity of ATF3, resulting in an impaired ATF3 function. ATF3 overexpression significantly reversed lipid accumulation decline and reduced inflammation after NR4A3 knockdown. Conclusion: The downregulation of NR4A3 alleviates MASLD by modulating ATF3, suggesting this may be a promising therapeutic target.
Chen, S.; Moorthy, A.; Yu, P. K.; Wang, J.; Liu, D.
Show abstract
With the increasing accessibility of single-cell RNA sequencing (scRNA-seq) data, cell-type-specific gene expression can be linked to complex traits through pseudo-bulk method, which considered aggregated gene expression from multiple cells of the same annotated cell type per individual and clearly shows the limitation of ignoring intra-individual cell-to-cell variability. Concurrently, pseudotime trajectory inference has gained popularity for its ability to capture continuous biological processes such as cell differentiation and lineage development, instead of individual discrete stages. It is natural to consider whether genetic effects for complex traits, such as individual level disease status, show a dynamic pattern along the inferred trajectories. In this study, we introduce a novel framework that models gene expression as a function of pseudotime along the inferred trajectories. We mapped expression quantitative trait loci (eQTL) effects in the cis-region as functional parameters, which we called "dynamic eQTLs", showing regulatory effects exerted by genetic variants change continuously along the cellular trajectory. For eQTLs of constant effects across pseudotime we leveraged external bulk-eQTL information to enhance the power. Furthermore, we employed significant, variable dynamic eQTLs as instrumental variables to infer causal relationships between gene expression and complex traits. To address challenges inherent to scRNA-seq data--such as sparsity and high variability--we incorporate an empirical likelihood-based inference method, which is non-parametric and self-normalized. Besides, genes associated with trajectory branchpoints may bring confounding, and we also proposed a causal mediation analysis framework to determine whether a gene plays a causal role for the disease directly and indirectly through driving cell fates. Applying our method to scRNA-seq data from human lung tissue of 114 samples (66 pulmonary fibrosis cases and 48 controls), along with meta-analyzed GWAS summary statistics for IPF from 3 studies, we identified pseudotime-dependent causal effects for IPF from genes implicated in the trajectory AT2 - translational AT2 - AT1, which is crucial in lung tissue repair and regeneration. We also found that 30 genes have a mediated effect through cell fates.
Oppong, A. E.; Louden, K.; HOLLOWAY, A.; ROSSI, L.; McDonnell, T. C. R.; Robinson, G. A.; ARULKUMARAN, N.; Manson, J. J.; Jury, E. C.
Show abstract
Haemophagocytic lymphohistiocytosis (HLH) is a rare, life-threatening hyperinflammatory syndrome characterised by uncontrolled immune activation. Reduced high- and low-density lipoprotein cholesterol and hypertriglyceridaemia are reported in HLH, suggesting lipid metabolism disturbances although in-depth serum metabolomic analysis is lacking in HLH. Here a lipid-focused NMR spectroscopy platform was used to define the serum metabolomic landscape of adults hospitalised with HLH compared to adults with sepsis (HLH-mimic) and rheumatic disease (potential HLH drivers/triggers), following surgical resection of solid organ cancer (non-infectious acute inflammation controls) and healthy controls (HCs). Serum metabolites distinguished HLH from HCs with high accuracy (>91.36%) using multiple machine learning models. The top classifying features included elevated apolipoprotein-B (ApoB)-containing low, intermediate, and very low-density lipoprotein particles; and lipoprotein remodelling towards triglyceride enrichment and cholesterol depletion. Differentially abundant metabolites in HLH compared to all control groups were enriched in pathways related to lipid metabolism including: 'Lipid particles composition', 'Plasma lipoprotein clearance', 'Plasma lipoprotein remodelling', 'Glucose homeostasis' and 'Amino acid metabolism'. Metabolomic results were validated using matched whole blood RNA-sequencing which identified differentially expressed genes enriched in metabolic modules associated with lipid, amino acid, and glucose metabolism, supporting a coordinated metabolic dysregulation in HLH from a transcriptomic to metabolomic level. Finally, twenty-seven metabolites including ApoB-containing, triglyceride-rich lipoproteins and saturated fatty acids distinguished HLH from all disease controls (AUC>0.70) either alone or combined as a metabolomic signature. Elevated ApoB and ApoB:ApoA1 ratio in HLH vs sepsis and HCs were validated by ELISA, supporting their utility as biomarkers to distinguish HLH from other hyperinflammatory syndromes.
Helgueta Romero, S.; Bonafina, A.; Olivie, N.; Coumans, B.; Nguyen, L.; Espuny Camacho, I.
Show abstract
The cerebellum is one of the most complex structures of the brain composed of a high diversity of GABAergic and glutamatergic neurons. Whereas cerebellar biogenesis has been extensively studied in the mouse, an in-depth characterization of genes and pathways involved in cerebellar specification and maturation in the humans remains overlooked. Here, we used human pluripotent stem cells (hPSC)-derived cerebellar organoids (CRBOs) to study the temporal biogenesis of neuronal subtypes. Our results show that CRBOs acquire caudal neural tube identity at an early stage followed by a time-dependent expression of mature cerebellar neuronal markers in vitro, mimicking human neurodevelopment. CRBOs show the generation of both cerebellar excitatory and inhibitory neurons and the expression of glial cell markers, suggesting the generation of a high variety of cerebellar cell types in vitro. Further, in vitro CRBOs show expression of cerebellar disease associated genes, such as those related to ataxia. Our results establish CRBOs as a valuable platform to explore the mechanisms of human cerebellar development and related disorders.
Li, D. J.
Show abstract
All cellular life forms fall under the three-domain classification of life, raising a fundamental evolutionary question: why does this classification feature three rather than two or four? To answer this question, a more general method, rather than the traditional one based on comparing small-subunit ribosomal RNAs, is required. The three-base periodicity in genomes is a common feature of both cellular life forms and viruses, which is species-specifically biased between amino acid biosynthetic families. Based on comparing such a common feature of all life forms, a global triangular diversification picture has been obtained, whose three angular regions correspond to the three domains, respectively. This mechanism of diversification of life attributes the evolutionary driving forces in diversification of the three domains of life to the biases between amino acid biosynthetic families. Notably, the same mechanism also applies to the contemporary diversification of SARS-CoV-2, whose reasonable results in turn corroborate the above explanation of primordial diversification of life and in addition shed light on the mechanism of speciation.
Sun, Y.; Yao, W.; Zhang, J.; Song, W.; Zhao, X.; Hao, C.; Chen, X.; Zeng, S.; Jia, S.; Yang, Y.; Chen, X.; Xiao, X.; Poo, M.-m.; Sun, Y.; Xu, B.; Zhang, T.
Show abstract
The organizational principles of natural neural networks could inspire the new architecture design of artificial neural networks (ANNs). Analysis of single-neuron connectomes of mouse brains revealed distinct profiles of three-node connectivity motifs in various cortical areas and hippocampal formation. A connectome-informed neural network algorithm ("CINA") was developed to incorporate natural connectivity motifs into ANN algorithms represented by recurrent neural network (RNN) and transformer-based large language model (LLM). We found that incorporation of the average profile of cortical motifs improved the RNNs performance in noise-resistant categorization and motor learning benchmark tasks, as compared with RNNs with random connectivity. Notably, incorporating cortex-specific motifs further elevated the RNNs performance in tasks related to the cortical function, and this effect was enhanced by artificially increasing the bias in the motif profile. Similar experimental results were verified on an LLM using Motif-Transformer for natural language question answering and brain-signal decoding tasks. Graph-theoretic analyses showed that incorporating natural motifs drove the emergence of modular and small-world properties in ANNs. Together, we demonstrated not only connectome-inspired optimization of ANN architecture but also functional significance of specific motif profiles in various cortices.
Grove, H.; Stenlokk, K. S. R.; Lien, S.; Gjuvsland, A. B.; Arnyasi, M.; van Son, M.; Kent, M.
Show abstract
Abstract The Duroc-derived reference genome Sscrofa11.1 has provided a critical foundation for pig genomics, providing a high-quality reference genome for accurate variant detection and comparative genomics but does not capture breed-specific variation. Here, we present a near-complete, gap-free genome assembly for the Landrace pig (Landrace_v1, GCA_963921485.1), spanning all 20 chromosomes and totaling 2.6 Gb, including 176 Mb of sequence absent from Sscrofa11.1. Comparative analyses with recently published high-quality pig genomes reveal a conserved centromere organization across breeds, accompanied by substantial variation in repeat composition and length, and identify a pig specific pattern of telomere variant repeats across eight pig breeds. The improved resolution of repetitive regions in Landrace_v1 enables more complete reconstruction of complex gene families, including olfactory receptors, and uncovers structural variation at the KIT proto-oncogene receptor tyrosine kinase locus not represented in the Duroc reference. Together, these findings highlight the limitations of single-reference genomes and demonstrate the value of breed-specific assemblies for capturing genomic diversity and improving downstream analyses.
Zhang, S.;Tyshkovskiy, A.;Ying, K.;Wang, S.;Gladyshev, V.
Show abstract
Alternative splicing exhibits significant changes during development and aging, affecting the composition and variance in the transcriptome. However, it is unclear whether and how age-associated splicing dysregulation leads to functional consequences. Here, an integrative analysis of transcriptome data across mouse and human tissues revealed that aging is characterized by systematic deterioration of the fidelity of RNA splicing, here termed splicing degeneration, a measure of functional alteration of reading frame and domain configuration of protein products. Genes with higher aging-associated splicing degeneration were more conserved and enriched for processes such as RNA metabolism and antigen presentation. By assessing alternative splicing events associated with functional deterioration, we quantified the degree of splicing degeneration. Its level increased with age but was alleviated following calorie restriction or rapamycin treatment, indicating that it can serve as a new molecular hallmark of aging. Mechanistically, through a comprehensive meta-data analysis, we discovered that splicing degeneration is associated with age-associated changes in specific splicing factors, which in turn showed a strong association with age-related transcriptome changes. Overall, our study demonstrates the intricate relationship between aging and genome-wide splicing degeneration, revealing a promising target for aging interventions acting to reverse splicing degeneration.
Picot, A.; Leboucher, M.; Helaine, C.; Talukdar, A.; Khalin, I.; Martinez de Lizarrondo, S.; Gauberti, M.; Nomenjanahary, M.; Goux, D.; Ho-Tin-Noe, B.; Vivien, D.; Bonnard, T.
Show abstract
Clot resistance to pharmacological thrombolysis remains a critical challenge in ischemic stroke (IS) management. Thrombus heterogeneity, particularly the presence of thrombolysis-resistant domains composed of dense fibrin and non-fibrin components, including neutrophil extracellular traps (NETs), significantly limits the efficacy of recombinant tissue-type plasminogen activator (r-tPA) and its variant, Tenecteplase (TNK). Consequently, novel therapeutic strategies are urgently required. Emerging evidence suggests that co-administration of deoxyribonuclease I (DNase I) with r-tPA can degrade DNA fibers and enhance clot lysis. In this study, we optimized a previously developed theranostic agent--iron oxide microparticles coated with polydopamine--by dual-grafting both r-tPA and DNase to target resistant thrombi. Using functional ultrasound imaging (fUS) during the acute phase of IS, we demonstrated accelerated reperfusion with this dual-functionalized platform in a r-tPA resistant IS model. Furthermore, MRI analysis confirmed a significant reduction in lesion volume at 24 hours, correlating with improved functional recovery five days post-ischemia.
Merino-Galan, L.;Hemenway, J.;Jagana, H.;Jackson, T.;Rajendran, A.;Khanna, A.;Ortiz-Espinosa, S.;Sarkar, S.;Kalia, V.;Pattwell, S.
Show abstract
The SARS-CoV-2 S1 protein is associated with immune cell activation and persistent neurological symptoms, yet the underlying mechanisms remain unclear, posing a major challenge in elucidating Long COVID pathophysiology. To investigate how circulating S1 contributes to long-term neurological alterations, we intravenously injected hACE2 mice with varying doses of S1 (5, 10, and 20 {micro}g) and observed temporally dysregulated systemic inflammatory responses accompanied by sub-acute CD4+ T cell infiltration into central limbic regions. This immune response induced mild sustained increases in cFos+ cells in the amygdala and mild neuroinflammation in the hippocampal CA1 region, resulting in both acute and long-term anxiety-like behaviors, while working memory remained unaffected. Together, these findings suggest that systemic S1 protein induces a sustained proinflammatory response that promotes lasting neurological alterations through immune-to-brain signaling pathways. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=136 SRC="FIGDIR/small/733010v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@10982caorg.highwire.dtl.DTLVardef@169a275org.highwire.dtl.DTLVardef@28e66aorg.highwire.dtl.DTLVardef@12f6c85_HPS_FORMAT_FIGEXP M_FIG C_FIG
Shengyi, Z.
Show abstract
Domains are the basic units of protein structure and function. Appropriate inter-domain organization is critical to enable cooperative execution of multiple related functions. It is thus a crucial step to determine the full-length structure of multi-domain proteins for the purpose of elucidating their functions and designing new drugs to regulate these functions. Existing structure prediction algorithms are generally better at solving the internal conformation of domains, rather than modeling the relative positions between domains. To address the challenge of accurately determining multi-domain protein conformations, we develop a single-sequence-based domain assembly algorithm called DDI_single. DDI_single directly extracts features from the amino acid sequence using the protein language model ESM-lb, and accurately predicts the interactions between residue pairs of structural domains through a novel gated cross-attention module, thus achieving the correct assembly of structural domains. With the knowledge of domain definition, DDI_single achieves more than 20% higher accuracy in the task of predicting the relative distances of residue pairs between domains than that of the single-sequence-based structure prediction algorithm trRosettaX_single. When assembling domains with known spatial conformations, DDI_single correctly assembles 74.4% of the samples in the test set (TM-score>0.5). When assembling domains with unknown spatial conformations, in cases where the internal spatial conformations of domains are correctly modeled, DDI_single correctly assembles 73.9% of the samples.
Ichikawa, Y.
Show abstract
Cross-population reversal of signed linkage disequilibrium (LD), or the "flip-flop" phenomenon, can arise when a tag SNP captures different extended haplotype backgrounds across populations. The MICA hepatocellular carcinoma susceptibility variant rs2596542 exemplifies this problem in the MHC, where signed LD reverses between Japanese and European populations but the relevant regulatory backgrounds are obscured by haplotypic complexity. We analyzed 7,303 biallelic SNVs surrounding rs2596542 across 26 populations using carrier-set topology classification followed by non-negative matrix factorization of carrier haplotypes. This identified two regulatory axes. Axis I, represented by components c4/c6, was population-stable and MICA-regulatory, with coherent MICA cis-eQTL enrichment and depletion for signed-LD reversal. Axis II, represented by component c5, was enriched for signed-LD reversal and showed an HLA-B{uparrow}/HLA-C{downarrow} expression signature with no MICA overlap across six GTEx tissues. In an independent Japanese HCC cohort (LIRI-JP, n = 122), Axis II-associated HLA-C downregulation remained after adjustment for clinical covariates, immune infiltration, and HLA-A expression. The previously proposed cross-population tag rs2244546 mapped to a population-stable component rather than Axis II. A parallel reanalysis of the COMT Val158Met flip-flop locus reproduced the signed-LD pattern reported by Lin et al1. and showed population-specific latent backgrounds among Val carriers. These results show that carrier-set topology combined with NMF can decompose composite marker alleles into functionally interpretable regulatory haplotype subspaces.
Koyama, S.; Nakao, T.; Choi, S. H.; Enzan, N.; Jurgens, S. J.; Ellinor, P. T.
Show abstract
Integrating human genetics into therapeutic discovery accelerates drug development. However, ancestral biases in historical cohorts have left critical functional variation largely uncharted. Here, we leverage the diverse NIH All of Us Research Program to conduct comprehensive common- and rare-variant association analyses for 624 quantitative traits across 369,655 ancestrally diverse individuals. We identified 6,181 genome-wide significant locus-trait associations (526 novel) and 416 gene-trait associations (105 novel) via rare-variant burden testing. By integrating fine-mapping with computational variant-effect predictors, we systematically prioritized rare, likely causal variants driving these signals. Jointly modeling common and rare variation with protein-class annotations significantly improved the identification of known drug targets compared to common-variant analysis alone. Notably, we identified NRG4 as a high-confidence candidate therapeutic target for preserving kidney function. Our findings demonstrate that characterization of rare and common variation across diverse populations enhances causal gene discovery and identifies novel, actionable therapeutic targets.
Shin, D.-H.; Jeon, J.; Joe, S.; Jeon, Y.; Yang, J. O.; Bhak, J.; Baek, S. A.; Byun, G.; Shin, E.-S.; Kwon, Y.; Choi, H.-J.; Kim, J.-H.; Haam, K.; Yoo, J.; Song, K. J.; Mok, J.; Jeon, S.; Jeong, H.; Bhak, J.
Show abstract
Here, we present the first graph-based Korean Pangenome Reference (K-PanRef), constructed from 14 healthy Korean individuals. K-PanRef comprises 13 high-quality diploid Korean genome assemblies (mean QV ~62.0) and KOREF1-G-TTAGGA, the first complete Korean reference genome. Integration of these assemblies generated a ~3.2-Gb pangenome graph containing ~39.3 million nodes and ~53.8 million edges, with the accumulation of common sequences (frequency [≥]10%) reaching a plateau. Additionally, K-PanRef contains ~4.3 million Korean-specific small variants and ~76.0 thousand Korean-specific SVs absent from the Chinese and human pangenome references, improving the representation of Korean genetic diversity relative to these references. To evaluate its utility for short-read-based SV analysis, we genotyped 75 whole-genome sequencing (WGS) samples, including 15 patients with early-onset myocardial infarction (MI). Although constructed entirely from healthy genomes, K-PanRef supported the identification of putative disease-relevant SVs in this exploratory application. K-PanRef-based genotyping identified ~95.6 thousand small variants and 820 SVs observed only in the early-onset MI samples. Among the early-onset MI-group SVs, 491 were absent from public databases, suggesting that they may represent previously unrecognized candidate variants related to early-onset MI. Of these, 164 SVs overlapped 134 genes, of which 89 had reported associations with 42 cardiovascular diseases or traits, including eight genes previously linked to MI. Together, these results establish K-PanRef as a valuable resource for representing Korean genetic diversity and enabling more comprehensive discovery of population-specific and novel putative disease-relevant variants from short-read sequencing data.
Xie, Z.; Xu, J.
Show abstract
MotivationFixed-backbone sequence design methods such as ProteinMPNN operate on backbone coordinates alone and cannot represent target side-chains at the binding interface. Their decoding algorithm also lacks a mechanism to balance binding affinity and folding stability or to improve selectivity against structurally similar off-targets. These gaps limit the computational design of protein binders with high affinity and specificity. ResultsWe present RedNet, a multiscale graph neural network that encodes side-chain information of the binding target. We further develop a contrastive decoding algorithm, motivated by the thermodynamic decomposition of binding free energy, that addresses two objectives: (1) balancing binding affinity and folding stability, and (2) improving selectivity against structurally similar off-targets. RedNet reaches 43% native sequence recovery on heterodimers, compared with 37% for ProteinMPNN and 33% for ESM-IF. With contrastive decoding, it matches native-sequence co-folding success (68%) on high-confidence AlphaFold3 targets, exceeding ProteinMPNN (59%) and ESM-IF (61%). On a new benchmark of structurally similar on-/off-target pairs, RedNet with contrastive decoding reaches 64.8% energetic selectivity, ahead of PiFold (55.6%), ProteinMPNN (53.7%), and ESM-IF (53.7%). AvailabilitySource code and datasets are released at https://github.com/zw2x/rednet_public. Contactjinbo.xu@gmail.com